Embedded Formulas Extraction
Identifieur interne : 000255 ( France/Analysis ); précédent : 000254; suivant : 000256Embedded Formulas Extraction
Auteurs : Afef Kacem [Tunisie] ; Abdel Belaïd [France] ; Mohamed Ben Ahmed [Tunisie]Source :
Descripteurs français
Abstract
A new approach for separating mathematics from usual text is presented. Contrary to the existing methods, it is more oriented toward the segmentation than the recognition, isolating the formulas outside and inside the text lines. The objective is to delimit a part of text which could disturb the OCR application, not yet trained for formula recognition and restructuring. The method is based on an adaptive segmentation working at two levels 1) A primary labelling identifies the more characteristic symbols; 2) A secondary labelling extends the context of the symbols for delimiting the formula inside the text.Experiments done on some commonly seen mathematical documents, show that our proposed method can achieve quite satisfactory rate making mathematical formulas extraction more feasible for real-world applications. The average rate of primary labelling of mathematical symbols is about 95.3% and their secondary labelling can improve the rate about 4%. Thus, about 95% of formulas are well extracted from images of documents printed with high quality
Url:
Affiliations:
- France, Tunisie
- Lorraine
- Nancy
- Institut national polytechnique de Lorraine, Université Nancy 2, Université de Lorraine
Links toward previous steps (curation, corpus...)
- to stream Hal, to step Corpus: 000046
- to stream Hal, to step Curation: 000046
- to stream Hal, to step Checkpoint: 000153
- to stream Main, to step Merge: 001D28
- to stream Main, to step Curation: 001C29
- to stream Main, to step Exploration: 001C29
- to stream France, to step Extraction: 000255
Links to Exploration step
Hal:inria-00099142Le document en format XML
<record><TEI><teiHeader><fileDesc><titleStmt><title xml:lang="en">Embedded Formulas Extraction</title>
<author><name sortKey="Kacem, Afef" sort="Kacem, Afef" uniqKey="Kacem A" first="Afef" last="Kacem">Afef Kacem</name>
<affiliation wicri:level="1"><hal:affiliation type="laboratory" xml:id="struct-421707" status="VALID"><orgName>Laboratoire RIADI-GDL [Manouba]</orgName>
<desc><address><addrLine>École Nationale des Sciences de l'Informatique (ENSI)Campus Universitaire de la Manouba,2010 Manouba, Tunisia</addrLine>
<country key="TN"></country>
</address>
<ref type="url">http://www.riadi.rnu.tn/</ref>
</desc>
<listRelation><relation active="#struct-301618" type="direct"></relation>
</listRelation>
<tutelles><tutelle active="#struct-301618" type="direct"><org type="institution" xml:id="struct-301618" status="INCOMING"><orgName>Ecole Nationale des Sciences de l'Informatique, Manouba, Tunisie</orgName>
<desc><address><country key="FR"></country>
</address>
</desc>
</org>
</tutelle>
</tutelles>
</hal:affiliation>
<country>Tunisie</country>
</affiliation>
</author>
<author><name sortKey="Belaid, Abdel" sort="Belaid, Abdel" uniqKey="Belaid A" first="Abdel" last="Belaïd">Abdel Belaïd</name>
<affiliation wicri:level="1"><hal:affiliation type="researchteam" xml:id="struct-2362" status="OLD"><orgName>READ</orgName>
<orgName type="acronym">READ</orgName>
<desc><address><country key="FR"></country>
</address>
</desc>
<listRelation><relation active="#struct-160" type="direct"></relation>
<relation name="UMR7503" active="#struct-441569" type="indirect"></relation>
<relation active="#struct-300009" type="indirect"></relation>
<relation active="#struct-300291" type="indirect"></relation>
<relation active="#struct-300292" type="indirect"></relation>
<relation active="#struct-300293" type="indirect"></relation>
</listRelation>
<tutelles><tutelle active="#struct-160" type="direct"><org type="laboratory" xml:id="struct-160" status="OLD"><orgName>Laboratoire Lorrain de Recherche en Informatique et ses Applications</orgName>
<orgName type="acronym">LORIA</orgName>
<desc><address><addrLine>Campus Scientifique BP 239 54506 Vandoeuvre-lès-Nancy Cedex</addrLine>
<country key="FR"></country>
</address>
<ref type="url">http://www.loria.fr</ref>
</desc>
<listRelation><relation name="UMR7503" active="#struct-441569" type="direct"></relation>
<relation active="#struct-300009" type="direct"></relation>
<relation active="#struct-300291" type="direct"></relation>
<relation active="#struct-300292" type="direct"></relation>
<relation active="#struct-300293" type="direct"></relation>
</listRelation>
</org>
</tutelle>
<tutelle name="UMR7503" active="#struct-441569" type="indirect"><org type="institution" xml:id="struct-441569" status="VALID"><idno type="IdRef">02636817X</idno>
<idno type="ISNI">0000000122597504</idno>
<orgName>Centre National de la Recherche Scientifique</orgName>
<orgName type="acronym">CNRS</orgName>
<date type="start">1939-10-19</date>
<desc><address><country key="FR"></country>
</address>
<ref type="url">http://www.cnrs.fr/</ref>
</desc>
</org>
</tutelle>
<tutelle active="#struct-300009" type="indirect"><org type="institution" xml:id="struct-300009" status="VALID"><orgName>Institut National de Recherche en Informatique et en Automatique</orgName>
<orgName type="acronym">Inria</orgName>
<desc><address><addrLine>Domaine de VoluceauRocquencourt - BP 10578153 Le Chesnay Cedex</addrLine>
<country key="FR"></country>
</address>
<ref type="url">http://www.inria.fr/en/</ref>
</desc>
</org>
</tutelle>
<tutelle active="#struct-300291" type="indirect"><org type="institution" xml:id="struct-300291" status="OLD"><orgName>Université Henri Poincaré - Nancy 1</orgName>
<orgName type="acronym">UHP</orgName>
<date type="end">2011-12-31</date>
<desc><address><addrLine>24-30 rue Lionnois, BP 60120, 54 003 NANCY cedex, France</addrLine>
<country key="FR"></country>
</address>
</desc>
</org>
</tutelle>
<tutelle active="#struct-300292" type="indirect"><org type="institution" xml:id="struct-300292" status="OLD"><orgName>Université Nancy 2</orgName>
<date type="end">2011-12-31</date>
<desc><address><addrLine>91 avenue de la Libération, BP 454, 54001 Nancy cedex</addrLine>
<country key="FR"></country>
</address>
</desc>
</org>
</tutelle>
<tutelle active="#struct-300293" type="indirect"><org type="institution" xml:id="struct-300293" status="OLD"><orgName>Institut National Polytechnique de Lorraine</orgName>
<orgName type="acronym">INPL</orgName>
<date type="end">2011-12-31</date>
<desc><address><country key="FR"></country>
</address>
</desc>
</org>
</tutelle>
</tutelles>
</hal:affiliation>
<country>France</country>
<placeName><settlement type="city">Nancy</settlement>
<region type="region" nuts="2">Lorraine</region>
</placeName>
<orgName type="university">Université Nancy 2</orgName>
<orgName type="institution" wicri:auto="newGroup">Université de Lorraine</orgName>
<placeName><settlement type="city">Nancy</settlement>
<region type="region" nuts="2">Lorraine</region>
</placeName>
<orgName type="university">Institut national polytechnique de Lorraine</orgName>
<orgName type="institution" wicri:auto="newGroup">Université de Lorraine</orgName>
</affiliation>
</author>
<author><name sortKey="Ben Ahmed, Mohamed" sort="Ben Ahmed, Mohamed" uniqKey="Ben Ahmed M" first="Mohamed" last="Ben Ahmed">Mohamed Ben Ahmed</name>
<affiliation wicri:level="1"><hal:affiliation type="laboratory" xml:id="struct-421707" status="VALID"><orgName>Laboratoire RIADI-GDL [Manouba]</orgName>
<desc><address><addrLine>École Nationale des Sciences de l'Informatique (ENSI)Campus Universitaire de la Manouba,2010 Manouba, Tunisia</addrLine>
<country key="TN"></country>
</address>
<ref type="url">http://www.riadi.rnu.tn/</ref>
</desc>
<listRelation><relation active="#struct-301618" type="direct"></relation>
</listRelation>
<tutelles><tutelle active="#struct-301618" type="direct"><org type="institution" xml:id="struct-301618" status="INCOMING"><orgName>Ecole Nationale des Sciences de l'Informatique, Manouba, Tunisie</orgName>
<desc><address><country key="FR"></country>
</address>
</desc>
</org>
</tutelle>
</tutelles>
</hal:affiliation>
<country>Tunisie</country>
</affiliation>
</author>
</titleStmt>
<publicationStmt><idno type="wicri:source">HAL</idno>
<idno type="RBID">Hal:inria-00099142</idno>
<idno type="halId">inria-00099142</idno>
<idno type="halUri">https://hal.inria.fr/inria-00099142</idno>
<idno type="url">https://hal.inria.fr/inria-00099142</idno>
<date when="2000-09">2000-09</date>
<idno type="wicri:Area/Hal/Corpus">000046</idno>
<idno type="wicri:Area/Hal/Curation">000046</idno>
<idno type="wicri:Area/Hal/Checkpoint">000153</idno>
<idno type="wicri:Area/Main/Merge">001D28</idno>
<idno type="wicri:Area/Main/Curation">001C29</idno>
<idno type="wicri:Area/Main/Exploration">001C29</idno>
<idno type="wicri:Area/France/Extraction">000255</idno>
</publicationStmt>
<sourceDesc><biblStruct><analytic><title xml:lang="en">Embedded Formulas Extraction</title>
<author><name sortKey="Kacem, Afef" sort="Kacem, Afef" uniqKey="Kacem A" first="Afef" last="Kacem">Afef Kacem</name>
<affiliation wicri:level="1"><hal:affiliation type="laboratory" xml:id="struct-421707" status="VALID"><orgName>Laboratoire RIADI-GDL [Manouba]</orgName>
<desc><address><addrLine>École Nationale des Sciences de l'Informatique (ENSI)Campus Universitaire de la Manouba,2010 Manouba, Tunisia</addrLine>
<country key="TN"></country>
</address>
<ref type="url">http://www.riadi.rnu.tn/</ref>
</desc>
<listRelation><relation active="#struct-301618" type="direct"></relation>
</listRelation>
<tutelles><tutelle active="#struct-301618" type="direct"><org type="institution" xml:id="struct-301618" status="INCOMING"><orgName>Ecole Nationale des Sciences de l'Informatique, Manouba, Tunisie</orgName>
<desc><address><country key="FR"></country>
</address>
</desc>
</org>
</tutelle>
</tutelles>
</hal:affiliation>
<country>Tunisie</country>
</affiliation>
</author>
<author><name sortKey="Belaid, Abdel" sort="Belaid, Abdel" uniqKey="Belaid A" first="Abdel" last="Belaïd">Abdel Belaïd</name>
<affiliation wicri:level="1"><hal:affiliation type="researchteam" xml:id="struct-2362" status="OLD"><orgName>READ</orgName>
<orgName type="acronym">READ</orgName>
<desc><address><country key="FR"></country>
</address>
</desc>
<listRelation><relation active="#struct-160" type="direct"></relation>
<relation name="UMR7503" active="#struct-441569" type="indirect"></relation>
<relation active="#struct-300009" type="indirect"></relation>
<relation active="#struct-300291" type="indirect"></relation>
<relation active="#struct-300292" type="indirect"></relation>
<relation active="#struct-300293" type="indirect"></relation>
</listRelation>
<tutelles><tutelle active="#struct-160" type="direct"><org type="laboratory" xml:id="struct-160" status="OLD"><orgName>Laboratoire Lorrain de Recherche en Informatique et ses Applications</orgName>
<orgName type="acronym">LORIA</orgName>
<desc><address><addrLine>Campus Scientifique BP 239 54506 Vandoeuvre-lès-Nancy Cedex</addrLine>
<country key="FR"></country>
</address>
<ref type="url">http://www.loria.fr</ref>
</desc>
<listRelation><relation name="UMR7503" active="#struct-441569" type="direct"></relation>
<relation active="#struct-300009" type="direct"></relation>
<relation active="#struct-300291" type="direct"></relation>
<relation active="#struct-300292" type="direct"></relation>
<relation active="#struct-300293" type="direct"></relation>
</listRelation>
</org>
</tutelle>
<tutelle name="UMR7503" active="#struct-441569" type="indirect"><org type="institution" xml:id="struct-441569" status="VALID"><idno type="IdRef">02636817X</idno>
<idno type="ISNI">0000000122597504</idno>
<orgName>Centre National de la Recherche Scientifique</orgName>
<orgName type="acronym">CNRS</orgName>
<date type="start">1939-10-19</date>
<desc><address><country key="FR"></country>
</address>
<ref type="url">http://www.cnrs.fr/</ref>
</desc>
</org>
</tutelle>
<tutelle active="#struct-300009" type="indirect"><org type="institution" xml:id="struct-300009" status="VALID"><orgName>Institut National de Recherche en Informatique et en Automatique</orgName>
<orgName type="acronym">Inria</orgName>
<desc><address><addrLine>Domaine de VoluceauRocquencourt - BP 10578153 Le Chesnay Cedex</addrLine>
<country key="FR"></country>
</address>
<ref type="url">http://www.inria.fr/en/</ref>
</desc>
</org>
</tutelle>
<tutelle active="#struct-300291" type="indirect"><org type="institution" xml:id="struct-300291" status="OLD"><orgName>Université Henri Poincaré - Nancy 1</orgName>
<orgName type="acronym">UHP</orgName>
<date type="end">2011-12-31</date>
<desc><address><addrLine>24-30 rue Lionnois, BP 60120, 54 003 NANCY cedex, France</addrLine>
<country key="FR"></country>
</address>
</desc>
</org>
</tutelle>
<tutelle active="#struct-300292" type="indirect"><org type="institution" xml:id="struct-300292" status="OLD"><orgName>Université Nancy 2</orgName>
<date type="end">2011-12-31</date>
<desc><address><addrLine>91 avenue de la Libération, BP 454, 54001 Nancy cedex</addrLine>
<country key="FR"></country>
</address>
</desc>
</org>
</tutelle>
<tutelle active="#struct-300293" type="indirect"><org type="institution" xml:id="struct-300293" status="OLD"><orgName>Institut National Polytechnique de Lorraine</orgName>
<orgName type="acronym">INPL</orgName>
<date type="end">2011-12-31</date>
<desc><address><country key="FR"></country>
</address>
</desc>
</org>
</tutelle>
</tutelles>
</hal:affiliation>
<country>France</country>
<placeName><settlement type="city">Nancy</settlement>
<region type="region" nuts="2">Lorraine</region>
</placeName>
<orgName type="university">Université Nancy 2</orgName>
<orgName type="institution" wicri:auto="newGroup">Université de Lorraine</orgName>
<placeName><settlement type="city">Nancy</settlement>
<region type="region" nuts="2">Lorraine</region>
</placeName>
<orgName type="university">Institut national polytechnique de Lorraine</orgName>
<orgName type="institution" wicri:auto="newGroup">Université de Lorraine</orgName>
</affiliation>
</author>
<author><name sortKey="Ben Ahmed, Mohamed" sort="Ben Ahmed, Mohamed" uniqKey="Ben Ahmed M" first="Mohamed" last="Ben Ahmed">Mohamed Ben Ahmed</name>
<affiliation wicri:level="1"><hal:affiliation type="laboratory" xml:id="struct-421707" status="VALID"><orgName>Laboratoire RIADI-GDL [Manouba]</orgName>
<desc><address><addrLine>École Nationale des Sciences de l'Informatique (ENSI)Campus Universitaire de la Manouba,2010 Manouba, Tunisia</addrLine>
<country key="TN"></country>
</address>
<ref type="url">http://www.riadi.rnu.tn/</ref>
</desc>
<listRelation><relation active="#struct-301618" type="direct"></relation>
</listRelation>
<tutelles><tutelle active="#struct-301618" type="direct"><org type="institution" xml:id="struct-301618" status="INCOMING"><orgName>Ecole Nationale des Sciences de l'Informatique, Manouba, Tunisie</orgName>
<desc><address><country key="FR"></country>
</address>
</desc>
</org>
</tutelle>
</tutelles>
</hal:affiliation>
<country>Tunisie</country>
</affiliation>
</author>
</analytic>
</biblStruct>
</sourceDesc>
</fileDesc>
<profileDesc><textClass><keywords scheme="mix" xml:lang="fr"><term>fuzzy logic mathematics segmentation</term>
<term>logique floue</term>
<term>segmentation de documents mathématiques</term>
</keywords>
</textClass>
</profileDesc>
</teiHeader>
<front><div type="abstract" xml:lang="en">A new approach for separating mathematics from usual text is presented. Contrary to the existing methods, it is more oriented toward the segmentation than the recognition, isolating the formulas outside and inside the text lines. The objective is to delimit a part of text which could disturb the OCR application, not yet trained for formula recognition and restructuring. The method is based on an adaptive segmentation working at two levels 1) A primary labelling identifies the more characteristic symbols; 2) A secondary labelling extends the context of the symbols for delimiting the formula inside the text.Experiments done on some commonly seen mathematical documents, show that our proposed method can achieve quite satisfactory rate making mathematical formulas extraction more feasible for real-world applications. The average rate of primary labelling of mathematical symbols is about 95.3% and their secondary labelling can improve the rate about 4%. Thus, about 95% of formulas are well extracted from images of documents printed with high quality</div>
</front>
</TEI>
<affiliations><list><country><li>France</li>
<li>Tunisie</li>
</country>
<region><li>Lorraine</li>
</region>
<settlement><li>Nancy</li>
</settlement>
<orgName><li>Institut national polytechnique de Lorraine</li>
<li>Université Nancy 2</li>
<li>Université de Lorraine</li>
</orgName>
</list>
<tree><country name="Tunisie"><noRegion><name sortKey="Kacem, Afef" sort="Kacem, Afef" uniqKey="Kacem A" first="Afef" last="Kacem">Afef Kacem</name>
</noRegion>
<name sortKey="Ben Ahmed, Mohamed" sort="Ben Ahmed, Mohamed" uniqKey="Ben Ahmed M" first="Mohamed" last="Ben Ahmed">Mohamed Ben Ahmed</name>
</country>
<country name="France"><region name="Lorraine"><name sortKey="Belaid, Abdel" sort="Belaid, Abdel" uniqKey="Belaid A" first="Abdel" last="Belaïd">Abdel Belaïd</name>
</region>
</country>
</tree>
</affiliations>
</record>
Pour manipuler ce document sous Unix (Dilib)
EXPLOR_STEP=$WICRI_ROOT/Ticri/CIDE/explor/OcrV1/Data/France/Analysis
HfdSelect -h $EXPLOR_STEP/biblio.hfd -nk 000255 | SxmlIndent | more
Ou
HfdSelect -h $EXPLOR_AREA/Data/France/Analysis/biblio.hfd -nk 000255 | SxmlIndent | more
Pour mettre un lien sur cette page dans le réseau Wicri
{{Explor lien |wiki= Ticri/CIDE |area= OcrV1 |flux= France |étape= Analysis |type= RBID |clé= Hal:inria-00099142 |texte= Embedded Formulas Extraction }}
This area was generated with Dilib version V0.6.32. |